PLOS Digital Health
● Public Library of Science (PLoS)
Preprints posted in the last 30 days, ranked by how well they match PLOS Digital Health's content profile, based on 106 papers previously published here. The average preprint has a 0.26% match score for this journal, so anything above that is already an above-average fit.
Dhaimade, P. A.; Henderson, R.
Show abstract
Multimodal large language models (MLLMs) are increasingly applied to image-based clinical reasoning, yet their diagnostic reliability in periodontal image interpretation, and the underlying source of their errors, remain poorly characterized. This study evaluated six architecturally distinct MLLMs (Claude Sonnet 4.5, GPT-5.0, Gemini 2.5, GLM-4.6, Sonar, and Grok 4.1) using 50 image-based multiple-choice questions drawn from the American Academy of Periodontology In-Service Examination, spanning clinical photographs, histopathology, radiographs, cardiac rhythm strips, and anatomical illustrations. A sequential two-phase experimental design was used: in Phase 1, each model independently described each image, selected an answer, and provided a supporting citation; in Phase 2, applied only to questions answered incorrectly, models were given an expert-validated visual description and asked to re-answer, allowing diagnostic improvement through visual correction to be measured directly. Expert ground truth for image content was established by a board-certified periodontist and independently validated by a second board-certified periodontist. Model outputs were classified using a dual-process error taxonomy adapted from Norman's model of diagnostic reasoning, distinguishing perceptual errors, arising from inaccurate visual feature extraction, from cognitive errors, arising from flawed reasoning despite accurate perception, with cognitive errors further subdivided into correctable and persistent subtypes, and additional categories capturing compound perceptual-cognitive failures and compensatory reasoning that overcame inaccurate perception. Diagnostic accuracy and error type distribution varied significantly across models and image modality. Correcting inaccurate visual descriptions in Phase 2 improved diagnostic accuracy for a subset of previously incorrect responses, indicating that a meaningful share of errors originated at the level of visual perception rather than clinical reasoning; conversely, a distinct subset of errors persisted despite accurate corrected visual input, indicating reasoning-level failures independent of perceptual accuracy. Some models also reached correct answers despite generating inaccurate image descriptions, reflecting compensatory reasoning resilient to perceptual error. These findings show that aggregate accuracy scores conflate mechanistically distinct failure modes, and that perceptual and cognitive errors carry different implications for how MLLMs might be safely deployed or improved for diagnostic image interpretation. The expert-guided visual correction framework introduced here provides a generalizable, mechanism-based approach to benchmarking multimodal AI diagnostic performance that extends beyond periodontics to other visually driven diagnostic domains in medicine. As MLLMs become increasingly accessible to clinicians, residents, and dental educators, distinguishing perceptual from cognitive failure is essential for guiding responsible clinical use, targeting model refinement, and informing AI-augmented dental education and competency assessment.
Al-Hebshi, S.; Khalifa, H.; Pham, T. D.
Show abstract
Background: Cone-beam computed tomography (CBCT) frequently captures the maxillary sinuses incidentally, and reliable automated detection of sinus abnormality is clinically relevant. Unlike most vision-language benchmarks in medical imaging, which pair images with pre-existing, human-authored clinical reports, findings text can also be generated directly by a large language model from the image itself--raising the question of how much diagnostic value such AI-derived text carries, and whether that value depends on independent verification. Multimodal artificial intelligence (AI) benchmarks risk overstating performance if the provenance of each input--image, raw AI-generated text, or radiologist-verified text--is not clearly separated and reported. Methods: We used 300 mid-sagittal CBCT slices from the MMDental dataset. ChatGPT generated findings text and a provisional normal/abnormal label for every slice (majority vote, three independent readings from the image alone); primary classification performance was assessed on this full, unfiltered set (n=300). A radiologist then independently reviewed each case's image together with ChatGPT's description, producing their own diagnosis; three cases were excluded as insufficient, yielding 297 confirmed cases. On this subset, every model was retrained and re-evaluated under identical 10-fold cross-validation on both the provisional ChatGPT-only labels ("pre") and the radiologist-confirmed labels ("post"), isolating the effect of label provenance from image or architecture. Eight vision architectures, seven language classifiers, and five VLMs were evaluated throughout; three generative models performed exploratory note-drafting. Findings: Raw ChatGPT-generated text produced the highest performance of any modality or condition: language models reached near-ceiling AUC (0.992 to 1.000, n=300), exceeding every vision model (AUC 0.799 to 0.880) and every VLM image-only probe (AUC 0.63 to 0.69). On the 297-case pre/post analysis, this advantage depended heavily on label source: language and text-derived VLM performance fell substantially from ChatGPT-only to radiologist-confirmed labels (e.g. BERT-base AUC 0.999 to 0.837), while vision-model performance was stable or modestly improved (e.g. DenseNet-121 0.867 to 0.891). The radiologist reclassified 62 of 297 cases (21%) relative to ChatGPT's provisional read, and a meaningful proportion of raw ChatGPT text was clinically uninterpretable or unsupported by the imaging. Interpretation: As shown here for the first time, raw, image-derived AI-generated text yields the highest apparent classification performance in this benchmark, but this reflects the text's alignment with its own self-generated labels rather than verified diagnostic content, and a substantial share of that text is not clinically explainable. Radiologist-confirmed text and labels give a lower but trustworthy estimate of true performance, on which convolutional neural network (CNN) vision models remain a stable, comparatively inexpensive baseline. Multimodal dental AI should report performance separately by modality and label provenance rather than pooling headline metrics.
Dasa, D.; Davies, P.
Show abstract
Objectives. To assess how digital inclusion factors and physical access barriers are associated with user trust in smartphone-based remote photoplethysmography (rPPG) hypertension screening, and to identify implications for digital health pol- icy, procurement and implementation in low-resource settings. Methods. Cross-sectional mixed-methods survey in five outpatient clinics in Kebbi State, northern Nigeria (N =287). Trust was measured using comfort, confidence and perceived usefulness Likert scales. Primary analyses used binary logistic models with HC3 robust standard errors; sensitivity analyses are reported in supplementary material. Free-text responses were thematically analysed. Results. Smartphone ownership was 51.2%; Transsion-brand devices comprised 56.5% of owners. Greater distance to a blood pressure facility was independently associated with lower perceived usefulness (OR 0.51, 95% CI 0.30-0.87; p=0.013) and lower comfort (OR 0.61, 0.37-0.98; p=0.042). Among owners, Transsion versus Samsung showed higher confidence odds (OR 3.82, 1.02-14.27; p=0.046). Qualitative themes supported the implementation interpretation: platform-fit and device speed requests among Transsion owners; connectivity and offline-first concerns among those with greater travel distance. No brand contrast achieved FDR-adjusted significance; brand findings are exploratory. Conclusions. Digital health policy and health technology assessment for smartphone-based screening should incorporate local device ecology, connectivity constraints, physical access burden and trust-calibration safeguards. Pre-implementation assessment of these factors is necessary for equitable and safe rPPG adoption in low-resource health systems.
Roberts, L.
Show abstract
Objective. Triage of rheumatology outpatient referrals is a high-volume administrative task that consumes senior specialist time without advancing patient care. The human triage system is only moderately accurate and reproducible. We assessed whether contemporary large language models (LLMs) are able to perform well enough to support automating this task in practice. In addition, the effects of different prompting techniques on triage accuracy and cost was assessed to help identify to optimal approach. Methods. Twenty referral scenarios spanning the urgency spectrum, based on real referrals were created by a certified Australian rheumatologist. Four rheumatologists triaged all cases independently and blinded, to produce a consensus reference standard. Twenty-three LLMs each triaged every referral into one of five urgency categories, three times (1380 outputs per condition). The experiment was run with a simple prompt and repeated with a advanced prompt supplying explicit triage expectations and worked examples. Results. All 2760 attempts returned valid categories. Under the simple prompt, performance separated into distinct tiers, larger models were more accurate (Spearman rho=0.42; P=.047) and accuracy tracked cost. Advanced prompting minimised between-model variance in accuracy 5.3-fold (0.014 to 0.003; Levene P=.01), abolished the size-accuracy association (rho=-0.05; P=.83) and removed the accuracy-cost relationship. Leading models matched expert consensus on most cases, within or above the range reported for human triage. Under-triage errors persisted with some LLMs. Conclusion. Contemporary LLMs categorise rheumatology referral urgency as well or better than published human triage systems. Advanced LLM prompting methods substitute for the reasoning capability of larger models, suggesting that LLM performance on this task may not require the most expensive models. The tools to automate this administrative task appear to already exist. Strong candidate LLMs that might serve a production ready solution have been identified.
Dick, M.; Madathil, S.; Patel, A.; Kapoor, H. S.; Sharma, M.; D'Souza, Z.; Hameed, S.; Abu-Samak, M.; Najirad, A.; Dwairi, D.; Radaideh, O.; Nicolau, B.
Show abstract
Objectives: Dentists prescribe approximately one in ten antibiotics worldwide, yet antimicrobial stewardship (AMS) remains underemphasized in dental education. Large language models (LLMs) may support AMS training, but their proficiency and clinical reasoning in this context remain unclear. We evaluated GPT-4o's accuracy and clinical reasoning on dental antibiotic prescribing questions, stratified by question difficulty. Methods: We assembled 125 multiple-choice questions on dental antibiotic prescribing from eight peer-reviewed studies (2017-2023). GPT-4o answered each question and generated a clinical justification. Accuracy was assessed against source-study answer keys and examined across difficulty quartiles. Justifications were evaluated using an adapted 12-axis human-evaluation framework assessing scientific consensus, extent and likelihood of harm, inappropriate and missing content, bias, and both correct and incorrect comprehension, retrieval, and reasoning. Prophylaxis-specific questions were analysed separately. Results: GPT-4o correctly answered 72% of questions. Accuracy remained relatively stable across difficulty quartiles (78%, 78%, 65%, 70%). Experts rated 95.4% of justifications positively across the 12 axes. Comprehension, retrieval, and reasoning each exceeded 96.2% positive ratings. Missing content was the main weakness (7.8%), and 7.1% of justifications showed a moderate-to-severe potential for harm. Performance on prophylaxis-specific questions (98.1%) exceeded non-prophylaxis questions (93.0%). Conclusions: GPT-4o demonstrated moderate-to-high proficiency and clinically defensible reasoning in dental antibiotic prescribing questions. However, residual risks indicate that it is not suitable for unsupervised clinical use but shows potential as a supervised AMS educational tool.
Mandich, A.; Koirala, S.; Westen, S.; Adhikari, S.; Acharya, A.; Shrestha, A.
Show abstract
Language discordance can impede community-based research and health communication where trained interpreters are limited. Although multimodal artificial intelligence systems can provide real-time spoken translation, performance with under-resourced languages during spontaneous field interactions remains poorly characterized. We evaluated ChatGPT-4o during bidirectional English-Nepali voice translation in a community setting near Dhulikhel Hospital, Nepal. In this cross-sectional field study, 30 primarily Nepali-speaking adults were recruited by convenience sampling. ChatGPT-4o mediated conversations using standardized English questions and spontaneous Nepali responses. A bilingual Nepali-English reviewer assessed 485 translated utterances using a 3-point accuracy scale and an inductively developed framework for translation and conversational deviations. Of 485 translations, 282 (58.1%) received the highest accuracy rating, 134 (27.6%) a moderate rating, and 69 (14.2%) the lowest. Mean accuracy was higher for English-to-Nepali than Nepali-to-English translation (2.63 {+/-} 0.53 vs 2.23 {+/-} 0.86); 63 of 69 low-accuracy translations (91.3%) occurred in the Nepali-to-English direction. Among 329 deviation tags, the most frequent were distortion of intended meaning (17.1%), overly formal or unnatural phrasing (14.7%), omission (14.2%), and addition of content (11.5%). Some fluent outputs substantially altered meaning or introduced information not expressed by the speaker. ChatGPT-4o demonstrated potential for real-time English-Nepali communication but also produced errors that could alter interpretation of participant responses. Accuracy was lower and more variable for Nepali-to-English translation; however, translation direction was confounded with input type because Nepali inputs were spontaneous and English inputs standardized, limiting conclusions about directional performance. These findings support cautious use for low-stakes conversational exchange and human verification when errors could affect research validity, clinical decisions, or participant understanding. As multimodal AI evolves, performance should be reevaluated across languages, real-world conditions, and model versions, with bilingual oversight and community partnership remaining central to responsible use.
kobayashi, v.; Baluyut, G. T. C.
Show abstract
Purpose Prevention and early detection of osteoporosis remains a global challenge, more so in regions like the Philippines where screening barriers exist. Chest x-rays meanwhile are relatively inexpensive, and more frequently done, and therefore can be used for opportunistic screening. This study aimed to develop a deep learning model for osteoporosis detection from chest x-rays using DXA as the gold standard. Methods A convolutional neural network called Osteo-AI was developed using 406 pairs of chest x-rays and DXA scans of Filipino patients aged 50 and above. With data augmentation, the training set expanded to 6,300 pairs. Gradient-weighted class activation mapping technique was applied to localize and identify patterns and areas in the chest x-ray images correlating with osteoporosis. Results Training data consisted of 369 female patients and 37 males. Ages of the patients ranged from 50 to 89 with a mean age of 63 years old. Initial testing yielded promising results, with Osteo-AI achieving a diagnostic accuracy of 85.71%, easily outperforming a benchmark of 33.33% Conclusion Our findings suggest the potential of Osteo-AI to enhance osteoporosis screening accessibility, aiding in early intervention to prevent fragility fractures. Further research involving larger datasets is warranted to refine and optimize the model, potentially improving detection accuracy and expanding its utility in global healthcare settings.
Xie, W.; Gupta, A.; Hossain, M. M.; Hasan, M.; Brage, S.; Forouhi, N.; Yadav, A.; Rajakaruna, V.; Gamage, M.; Mahmood, S.; Rajendra, P.; Jha, V.; Kasturiratne, A.; Katulanda, P.; Khawaja, K. I.; Mridha, M. K.; Hersch, F.; Anjana, R. M.; Chambers, J.; Goon, I. Y.
Show abstract
Abstract Background: A critical challenge for large-scale multi-country population health studies is the ability to collect consistent data across many sites and time periods and ensure that the data collected are valid and comparable. The use of mobile digital devices coupled with data collection platforms can address this challenge. We developed a fit-for-purpose digital data collection platform for the South Asia Biobank study. Objective: To describe the process by which a digital platform was designed, developed and deployed across four countries in South Asia; to demonstrate how the platform enabled field research teams located across these countries to collect non-communicable diseases epidemiological data consistently. Methods: A user-centred design approach was employed for the development of the digital platform to address the dynamic nature of study requirements. This approach uses 5-step iterative loops that, with each iteration, produce a usable prototype version of the software that was then tested by potential users of the platform. Qualitative interviews and quantitative system usability assessments were conducted, and findings utilised as input for the start of the next iterative loop. The process was completed when a working version of the software was developed for the use in the study. Results: Over the course of four iterative loops, the platform was progressively built and tested to ensure its functionality met the requirements of the study. Detailed feedback was collected from key stakeholders and incorporated into the platform with each new version of the applications. The platform leverages advances in mobile and medical device technology along with software integration capabilities to enable efficient and consistent data collection, along with the ability to review data quality and make improvements to the data collection process in real-time. The successful deployment of the data platform has enabled collection of comprehensive baseline data from 205,536 participants in four South Asian countries. Conclusions: Using user-centred design principles, it is possible to develop and deploy a comprehensive digital surveillance data management platform that allows consistent and high-quality data collection in population health studies in remote settings. To the best of our knowledge, this is the first platform that enables the integrated capture of health assessment data from a wide variety of medical equipment that is tailored for deployment in a range of LMIC settings.
Rony, A. R.; Nahin, K. S. A.; Islam, T.; Asha, A. S.; Hossen, A.
Show abstract
Caesarean section in Bangladesh reached 51.8% of deliveries in 2025, and elective caesarean, meaning caesarean before labour began, reached 31.6%. Risk models built on national household surveys are increasingly proposed for pointing audit toward places where scheduled surgery is outrunning clinical need, but they are usually validated in ways that flatter them. Using the 2025 Bangladesh Multiple Indicator Cluster Survey, we developed four models on 9,538 women (logistic regression, elastic net, random forest, gradient boosting) and ran the same procedure under three validation designs: random five-fold cross-validation; five-fold cross-validation grouped by sampling cluster; and leave-one-division-out cross-validation. We also tested transfer between the 2019 and 2025 rounds and audited subgroup calibration. No model improved on logistic regression by a margin worth acting on: the area under the receiver operating characteristic curve ranged from 0.724 to 0.736 under cluster-grouped validation, a spread of 0.012. Validation design mattered far more than the algorithm. Grouping folds by sampling cluster changed discrimination by at most 0.0004, this survey contributing a median of 3 eligible women per enumeration area. Withholding a whole division cost 0.044 to 0.060, more than 100 times as much, and still cost 0.033 to 0.056 after the strongest predictor, an outcome-derived district rate, was removed from every model. A model fitted to 2019 data lost 0.083 when applied to 2025, and the two rounds agreed only moderately on which predictors mattered (Spearman rank correlation 0.61). Calibration held in every wealth quintile, both residence categories and seven of eight divisions; Sylhet was the exception. Elective caesarean is predictable from routine survey items, but that predictability is local. Cross-validation, including cluster-aware cross-validation, does not measure what a model would do in a district it has never seen; a geographic holdout is the cheapest design that does.
Christiansen, A.; Page, R.
Show abstract
TikTok has become a significant source of health information, and concern has grown about AI-generated content (henceforth, 'AIGC') as a vehicle for health misinformation. Where AIGC presents realistic-appearing people giving health advice, disclosure labels are the viewer's only reliable cue that what they are watching is synthetic. This research letter compares AI label metadata across 128,016 mental health-related TikTok videos and 4,924 videos from a network of 50 profiles posting exclusively AI-generated mental health content to evaluate how much content reaches audiences undisclosed. In a keywords-based collection, fewer than a percent of TikTok videos about mental health carried an AI label, but in profiles containing purely AI-generated content, just over 9 in 10 videos (90.23%) were neither labelled by the creator nor identified by TikTok's automatic detection. Additionally, in the keyword collection, automatic detection produced the majority of labels, while in confirmed AI-generated content from 50 profiles, it accounted for just three of the 481 labelled videos. These findings highlight the challenging landscape of AI disclosure and labelling and raise questions about where automatic detection is failing.
Davies, J.; Biondi, A.; Viana, P. F.; Ampe, L.; Schreiber, J.; Richardson, M. P.
Show abstract
Seizure diaries are one of the most useful sources of information in the management of epilepsy, however patient engagement with them can be sporadic. Sustained participation with seizure diaries affects the completeness and reliability of self-reported data, so it is vital to be able to measure engagement. To facilitate this, we create a multidimensional engagement metric with which to characterize how patients interact with their seizure diary. We utilise data from the Helpilepsy, a seizure diary application, common features found in application engagement metrics in business settings, and well understood clinical features to do this. Clustering is then performed to isolate different user groups based on how engaged they are, and these groups are studied to understand what drives the differences in engagement. We found three groups emerge from the clustering: low, medium and highly engaged users. Investigating these groups further, we put together a ``profile" for highly-engaged users. We find that they tend to be older at the point of diagnosis, and have had epilepsy for longer than the other users. We also find they tend to have had more medications, have higher doses of common anti-seizure medications, and they have more medications typically given to those with refractory epilepsy. The implications for e-diary design are that more attention should be given to those newer to epilepsy in the onboarding phase. Also, engagement is not necessarily based on just the upload of seizures, with other features of an e-diary being important to be filled in.
Chepngeno, J.; Rosen, R. K.; Lantini, R.; Garbern, S. C.; Salvatory, M.; Rameck, R.; Dhalla, F.; Yu, D.; Sharma, V.; Duggan, C.; Manji, K. P.; Levine, A. C.
Show abstract
Background: In two large studies conducted in Bangladesh, our recently developed artificial intelligence (AI)-based models for assessing dehydration severity in children under five years (DHAKA models) and patients over age five (NIRUDAK models) were significantly more accurate and reliable than the WHO IMCI and IMAI guidelines for diarrhea management. We incorporated these models into a novel mobile health (mHealth) clinical decision support tool (CDST), called FluidCalc, with the potential to improve acute diarrhea management by frontline health workers worldwide. Our objective was to assess the barriers and facilitators to uptake and use of our mHealth CDST in both a low-resource setting (Tanzania) and high-resource setting (United States (US)) among healthcare providers and stakeholders. Methods: Qualitative data were collected through focus group discussions (FGDs) with healthcare providers and in-depth interviews (IDIs) with stakeholders and policymakers from February - July 2025 in Tanzania and February - March 2026 in the US. The Consolidated Framework for Implementation Research (CFIR) was used to guide discussions and elicit participant feedback. Audio recordings were transcribed and translated from Swahili to English where applicable, and data were analyzed using framework matrix analysis. Results: 35 providers from different cadres participated in FGDs, and 13 stakeholders participated in IDIs. Facilitators to implementation included FluidCalc's simplicity, ease of use, and offline functionality. Participants reported that the app could streamline clinical workflows, promote adherence to diarrhea management guidelines, facilitate task shifting, support antibiotic stewardship, and reduce errors in fluid rehydration calculations. FluidCalc was also viewed as a valuable teaching tool, and for supporting less experienced healthcare providers and trainees, and as useful during diarrheal disease outbreaks. Perceived barriers included the need for reliable digital infrastructure, including access to mobile devices, internet connectivity, and dependable electricity and lengthy institutional approval processes. Endorsement and approval from the Ministry of Health and health facility leadership were perceived as essential for successful implementation. Conclusion: Healthcare providers and stakeholders believe FluidCalc has the potential to improve care for patients with acute diarrhea in both high- and low-resource settings. Addressing identified barriers and ensuring reliable digital health infrastructure are needed to support effective integration into patient care.
Steitz, B. D.; Ogunsan, O. O.; Ancker, J. S.; Carlson, B. R.; Gaynor, L. S.; Higashi, R. T.; Morrow, E. L.; Reese, T. J.; Romano, R. R.; Stern, S.; Turer, R. W.; Rosenbloom, S. T.; Wright, A.
Show abstract
Objectives: Characterizing patient portal message content at scale can help target efforts to manage administrative work. We developed and validated a large language model (LLM) pipeline for multi-label classification of messages using an expert-derived topic taxonomy, then characterized topic distribution across a two-year corpus. Materials and Methods: We studied all medical advice request messages sent to ambulatory clinicians at an academic medical center from 2024-2025. We convened an expert panel that derived an 11-category taxonomy through a modified Delphi process. Two annotators labeled 750 randomly selected messages (Cohen kappa 0.80), holding out 500 for evaluation. The pipeline used GPT-4o-mini in a zero-shot prompt. On the held-out set, we measured micro- and macro-averaged precision, recall, and F1, and label stability across runs. We then characterized topic distribution and co-occurrence across the corpus. Results: The pipeline achieved micro- and macro-averaged F1 of 0.89 and 0.86. Labels were identical across runs for 93.6% of messages. Across 2.4 million messages, content concentrated on a few topics. The two most common topics, Problems & Management and Medications & Prescriptions, were present in 67.9% of messages, and the four most common in 93.9%. 51.7% of messages addressed multiple topics. Discussion and Conclusion: The pipeline classified patient message topics accurately and stably across millions of messages. Message content was concentrated within a small number of topics, highlighting opportunities for targeted interventions and enabling more efficient triage, routing, and patient-facing support.
Wojcik, S.; Rulkiewicz, A.; Domienik-Karłowicz, J.
Show abstract
Large language models perform well on medical examinations, but users routinely challenge their answers and invoke professional roles, and it is unclear what a system does when a medical credential and a stated task-specific accuracy point in opposite directions. In a factorial experiment on 480 items from four Polish specialty examination sets and three consumer large language model systems (ChatGPT, Claude, Gemini), each item and system received eleven independent conversations. Conditions crossed attributed source role (medical student, experienced specialist), stated prior accuracy on similar questions (2/10, 8/10) and suggestion correctness. The primary outcome was adoption of a prespecified incorrect option when the baseline answer matched the official key, comparing a specialist described as 2/10 with a student described as 8/10. Baseline agreement with the key was 87.2% across 15,683 analyzable conversations. The incorrect option was adopted more often from the specialist described as 2/10 than from the student described as 8/10 (10.2% vs. 7.6%; adjusted risk difference +2.82 percentage points, 95% CI +0.65 to +4.99). Estimates varied across the three systems and only one system-specific interval excluded zero. In a prespecified exploratory analysis with a shared eligibility rule, correct suggestions were adopted far more often than incorrect ones (risk difference +35.7 percentage points, 95% CI +30.8 to +40.7), indicating selective rather than indiscriminate compliance. An incorrect suggestion from a specialist with low stated accuracy was therefore slightly more influential than the same suggestion from a student with high stated accuracy, although the difference was modest and varied across systems. Agreement reached only after a user has disclosed a preferred answer should not automatically be treated as an independent second opinion, and medical large language model systems should be evaluated on how they revise answers after such disclosure, not solely on initial accuracy.
Champeaux, S. A.; Booth, J.; Brown, A.; Sebire, N. J.; Drobnjak, I.; Bowyer, S.
Show abstract
Background: Machine learning models leveraging electronic health records (EHRs) can support earlier detection of sepsis in intensive care units (ICUs). However, their clinical utility depends on reproducibility across institutions and patient populations. Building on a published pipeline from the Children's Hospital of Philadelphia (CHOP), this study examines how a neonatal sepsis prediction framework performs and can be adapted to a range of intensive care environments, paediatric, cardiac, and neonatal, at Great Ormond Street Hospital (GOSH). Methods: We extracted de-identified ICU EHR data from GOSH and applied feature derivation, unit harmonisation, and temporal sampling to align with the CHOP dataset used by Masino et al. (2019). Seven classifiers were first evaluated using CHOP-trained weights to characterise cross-domain behaviour and then retrained on local data to assess recoverability and site-specific adaptation. Model discrimination was summarised by AUC and F1, and learning curves were used to explore sample efficiency and bias-variance dynamics. Results: Models achieved strong discrimination on the CHOP neonatal cohort but demonstrated reduced performance when transferred to the mixed GOSH ICU population, reflecting anticipated domain and population shift. Retraining on GOSH data restored discrimination (AUC range 0.69-0.86), with Gradient Boosting (AUC 0.86 vs AUC 0.87 at CHOP) and KNN (AUC 0.80 vs AUC 0.79 at CHOP) models performing comparably to their CHOP benchmarks. DeLong's test confirmed statistically significant gains across all classifiers (p < 0.001). Conclusion: ICU cohort and baseline demographic differences between CHOP and GOSH introduced domain shift that limited direct model transfer. Elements of the original preprocessing pipeline could not be reproduced, further constraining transportability. Yet, retraining on local data restored high discrimination, showing that the modelling framework remains robust when re-estimated in new settings. These results highlight local adaptation as a practical route to recover performance and support safe, generalisable deployment of clinical prediction models in mixed clinical environments.
Ho, L. Y.-L.; Wong, K. C.-Y.; Cheng, L. W.-K.; Wan, A. T.-Y.; She, C. H.; Tsang, K. L. V.; So, H.-C.; Tsui, S. K.-W.
Show abstract
The rising prevalence of autism spectrum disorder (ASD) strains clinical infrastructure. Gold-standard tools like ADOS-2 face high costs, specialized training requirements, and extensive waitlists, delaying diagnosis and intervention. While eye-tracking offers a promising digital biomarker, existing tools lack scalable community deployment due to hardware costs and operational constraints. Here, we introduce the WISE-Screen framework, a smartphone-based real-time architecture for autonomous ASD Screening and multidimensional phenotypic profiling, evaluating its conceptual feasibility across a development-tally diverse age range. Two machine learning pipelines processed smartphone-captured eye-gaze data: (1) a Scanpath-based (SP) pipeline utilizing saliency maps and engineered scanpath features across 34 stimuli to estimate ASD-typical gaze probabilities, and (2) a Domain-task-based (DT) pipeline evaluating responses to 17 specialized tasks across four phenotypic domains (social, emotional, sensory, executive). Models were evaluated using leave-one-out cross-validation on 35 participants (16 ASD, 19 Non-ASD, ages 2.5-17) with ADOS-2 confirmed status. Compared to a baseline demographic model (ROC-AUC = 0.82; 95% CI: 0.68-0.96), performance improved using SP model (ROC-AUC = 0.90; 95% CI: 0.78-1.00) and DT model (ROC-AUC = 0.88; 95% CI: 0.75-1.00), with the integrated model reaching a peak ROC-AUC of 0.91 (95% CI: 0.80-1.00). Age- and sex-residualized models maintained an adjusted ROC-AUC of 0.74 (95% CI:0.57-0.92), with sensory, social and emotional domains showing the strongest association. WISE-Screen offers a scalable, automated adjunct to traditional protocols, providing accessible digital phenotyping to overcome systemic ASD screening barriers, though further evaluation in larger cohorts is warranted.
Jo, A. A.
Show abstract
Maternal healthcare prediction systems often suffer from algorithmic biases due to socio-economic disparities and imbalanced datasets, limiting their effectiveness for equitable healthcare policymaking. This paper introduces MaternaAI, a fairness-aware and explainable learning framework designed to enhance maternal healthcare predictions in Kerala, India. The framework focuses on three critical health indicators:(1) Tetanus Toxoid (TT) booster uptake,(2) immunization coverage rates, and (3) the percentage of pregnant women completing four or more Antenatal Care (ANC) visits. To address fairness, we propose Adaptive Equity Score Optimization (AESO), a novel optimization algorithm that dynamically integrates fairness constraints into model training. AESO is model-agnostic and adapts group equity weights in response to real-time disparities. We integrate SHAP, LIME, and feature permutation techniques for explainability, enabling transparent global and local interpretation. Empirical results demonstrate that MaternaAI significantly improves fairness metrics and model accuracy across diverse machine learning and deep learning models, offering interpretable and equitable decision support for public health stakeholders.
Kalla, M.; Bray, S. C.; Schadewaldt, V.; Krishnasamy, M.; Whittle, J. R.; Chapman, W.; Huckvale, K.; Burns, K.; Capurro, D.; Layton, M. J.; Thomas, J.; Lourenco, R. D. A.; Andrew, D.; McAlpine, H.; Dhillon, R. S.; Cain, S.; Rosenthal, M.; Drummond, K. J.
Show abstract
Patients with a brain tumour receive evidence-based clinical care in Australia but a focus on supportive care, including social connection, is often deficient. Digital health platforms hold promise to support these patients and their carers. Existing platforms often lack end-user co-design, evidence-based development and rigorous evaluation. Recognising this unmet need, we co-designed Brain Tumours Online, a digital supportive care platform to streamline access to educational resources, symptom management tools, and peer support for patients, carers, and healthcare professionals. In this article, we present our evaluation approach for Brain Tumours Online to advance methodological thinking in the evaluation of multi-faceted, co-designed digital health platforms. In contrast to standardised procedures in clinical trials, digital health interventions such as supportive care platforms are more complex due to their interactive nature, no prescriptive protocols for usage and the dynamic content of web-based information. Thus, traditional evaluation approaches often fall short in evaluating such multi-faceted digital health supportive care platforms. To address these challenges, we developed a bespoke, logic-modelling based evaluation approach to assess the usability, engagement, impact, and economic value of our platform. Our pragmatic but rigourous evaluation approach required the adaptation of existing evaluation frameworks, subject-matter, and lived experience expert knowledge. Our implementation science and co-design approach are shared in different papers. Our study outcomes will also be shared in a separate paper. In the current paper, we share our approach to the evaluation of Brain Tumours Online and provide insights that may be of value for other researchers interested in the nuances of trialing multi-faceted digital health supportive care platforms.
Srivastava, D. K.; Gupta, S.; Yadav, N.
Show abstract
Background: Evaluation of public health surveillance systems is a programmatic obligation but has largely been conducted as a periodic, externally commissioned activity requiring dedicated resources and additional data collection. India's Integrated Disease Surveillance Program (IDSP) generates continuous outbreak data through weekly reports but lacks a routine, embedded performance evaluation mechanism. This study assessed the quality of IDSP outbreak detection and response across multiple surveillance attributes and developed a weighted composite performance scoring framework using only routine program data. Methods: A cross-sectional evaluation study was conducted across 38 districts of Bihar using secondary data from IDSP Central Surveillance Unit weekly outbreak reports for 2016 - 2018 (n=559 outbreaks). Six surveillance quality attributes were assessed - timeliness, completeness, representativeness, relative sensitivity, acceptability and flexibility. A weighted composite performance scoring scale was developed using expert opinion-derived attribute weightages (n=25 experts). District-level scores were computed and scaled to 100. Results: Timeliness was the poorest-performing attribute, with fewer than 15% of outbreaks notified within 48 hours across all three years. Private sector participation was entirely absent - the acceptability score was 0 across all 38 districts for all three years. Completeness was the strongest attribute, exceeding 95% in all years. The mean composite score remained consistently low (23 - 27 out of 100) with widening inter-district disparity over time. Four districts (10.5%) scored 0 in all three years. Conclusions: This study presents a dynamic, routine-data-based composite performance evaluation framework for IDSP outbreak detection and response. The modular, configurable framework functions at any administrative level (from block to national) and is compatible with digital health information platforms, enabling continuous, embedded performance monitoring without additional data collection. The framework has been registered as an Intellectual Property with the Government of India. Keywords: Disease surveillance; IDSP; IDSR; performance evaluation; composite score; outbreak detection; timeliness; completeness; relative sensitivity; digital health
Edmond, E. C.; Dreyer, A. J.; Winston, A.; Khoo, S. H.; Joska, J.; Nightingale, S.
Show abstract
Background Computerised cognitive testing may address the global challenge in identifying cognitive changes in people living with HIV scalably and affordably. We assessed a computerised battery (CB) of cognitive tests, in a prospective cohort (CONNECT) of people with HIV in a low-income peri-urban area of Cape Town, South Africa during a national programmatic switch from efavirenz- to dolutegravir-based antiretroviral therapy (ART). Methods We recruited 170 people with HIV and 91 people without HIV (controls) (140[82%] and 41[45%] followed up). The CB and gold-standard pen&paper cognitive testing (P&P) were performed at both timepoints. Technology familiarity/use questionnaire data were also collected. We compared performance in detecting lower group-level cognitive performance associated with efavirenz treatment. Furthermore, the CB was compared to P&P in classifying individuals with low cognitive performance, correlation of global test scores and domain-level scores between batteries, and practice effects between timepoints. Exploratory principal component analysis was also performed. Results People with HIV on efavirenz at baseline had lower performance on the computerised battery than controls, {Delta}T=2.6, p=0.0047. This difference was lost after switching to dolutegravir-based ART at follow-up. CB and P&P global T were moderately correlated (R2=0.203, p<0.001), and the CB performed moderately in classification of low cognitive performance against the gold standard (AUC 0.70, sensitivity 0.52, specificity 0.76, PPV 0.40, and NPV 0.84). Selecting the first three principal components improved both classification of low cognitive performance (AUC 0.77) and correlation strength with P&P global T (R2=0.3, p<0.001). The CB did not show practice effects. Most participants owned a mobile phone (95%, 85.9% of these smartphones). Performance was better in smartphone owners ({Delta}T=1.8) and computer owners (23%, {Delta}T=1.8). Conclusions Delivering computerised cognitive testing was feasible in this low-income southern African setting. The CB showed reasonable construct validity (detecting known lower cognitive performance associated with efavirenz-ART) and may detect broad cognitive characteristics such as processing speed and accuracy. However, correlation of CB results with gold standard P&P testing was low-moderate and may limit its applicability as a diagnostic tool. This might be improved by including a wider range of cognitive domains tested in the CB, or data driven analysis. Brief CBs may fulfil an initial screening role, followed by more detailed clinical assessment.